Search CORE

Systematic identification of conserved motif modules in the human genome

Author: A Subramanian
A Visel
AL Donner
B Ren
CE Lawrence
CS Shashikant
DC King
DJ Galas
DS Johnson
E Eden
E Wingender
EH Davidson
EH Margulies
G Grahne
G Robertson
GD Stormo
GG Loots
GG Prefontaine
Haiyan Hu
HJ Bussemaker
J Han
J Hu
JC Knight
JD Hughes
KH Lee
L Narlikar
Lin Hou
M Blanchette
M Blanchette
M Brudno
M Fried
M Gupta
MA Eid
Minghua Deng
MM Garner
Naifang Su
NB La Thangue
OV Kel-Margoulis
PR Stabach
Q Zhou
S Sinha
SA Sholl
TL Bailey
WW Wasserman
WW Wasserman
X Cai
X Li
X Zhang
Xiaohui Cai
Xiaoman Li
Publication venue: BioMed Central
Publication date: 01/01/2010
Field of study

Abstract Background The identification of motif modules, groups of multiple motifs frequently occurring in DNA sequences, is one of the most important tasks necessary for annotating the human genome. Current approaches to identifying motif modules are often restricted to searches within promoter regions or rely on multiple genome alignments. However, the promoter regions only account for a limited number of locations where transcription factor binding sites can occur, and multiple genome alignments often cannot align binding sites with their true counterparts because of the short and degenerative nature of these transcription factor binding sites. Results To identify motif modules systematically, we developed a computational method for the entire non-coding regions around human genes that does not rely upon the use of multiple genome alignments. First, we selected orthologous DNA blocks approximately 1-kilobase in length based on discontiguous sequence similarity. Next, we scanned the conserved segments in these blocks using known motifs in the TRANSFAC database. Finally, a frequent pattern mining technique was applied to identify motif modules within these blocks. In total, with a false discovery rate cutoff of 0.05, we predicted 3,161,839 motif modules, 90.8% of which are supported by various forms of functional evidence. Compared with experimental data from 14 ChIP-seq experiments, on average, our methods predicted 69.6% of the ChIP-seq peaks with TFBSs of multiple TFs. Our findings also show that many motif modules have distance preference and order preference among the motifs, which further supports the functionality of these predictions. Conclusions Our work provides a large-scale prediction of motif modules in mammals, which will facilitate the understanding of gene regulation in a systematic way.</p

eScholarship - University of California

University of Central Florida (UCF): STARS (Showcase of Text, Archives, Research & Scholarship)

Identification of gene co-regulatory modules and associated cis-elements involved in degenerative heart disease

Author: A Subramanian
AI Su
Arkady M Pertsov
AS Barth
BJ Wilkins
C Danko
C Kioussi
Charles G Danko
DW Jeong
E Segal
F Tan
F Wittchen
G Dennis
H Rindt
H Wakaguri
J Hwang
J Tian
J Wang
JA Towbin
JD Barrans
JL Hall
KA Dellow
LA Megeney
M Flesch
M Gupta
MA Beer
MB Eisen
MM Kittleson
MS Parmacek
MS Parmacek
OV Kel-Margoulis
PK Bhavsar
R Bassel-Duby
R Development Core Team
R Edgar
R Gentleman
R Grzeskowiak
RCG Holland
S Malik
T Sugimoto
TH Christensen
TJP Hubbard
VR Iyer
WE Johnson
X Xie
Publication venue: BioMed Central
Publication date: 01/01/2009
Field of study

Abstract Background Cardiomyopathies, degenerative diseases of cardiac muscle, are among the leading causes of death in the developed world. Microarray studies of cardiomyopathies have identified up to several hundred genes that significantly alter their expression patterns as the disease progresses. However, the regulatory mechanisms driving these changes, in particular the networks of transcription factors involved, remain poorly understood. Our goals are (A) to identify modules of co-regulated genes that undergo similar changes in expression in various types of cardiomyopathies, and (B) to reveal the specific pattern of transcription factor binding sites, <it>cis</it>-elements, in the proximal promoter region of genes comprising such modules. Methods We analyzed 149 microarray samples from human hypertrophic and dilated cardiomyopathies of various etiologies. Hierarchical clustering and Gene Ontology annotations were applied to identify modules enriched in genes with highly correlated expression and a similar physiological function. To discover motifs that may underly changes in expression, we used the promoter regions for genes in three of the most interesting modules as input to motif discovery algorithms. The resulting motifs were used to construct a probabilistic model predictive of changes in expression across different cardiomyopathies. Results We found that three modules with the highest degree of functional enrichment contain genes involved in myocardial contraction (n = 9), energy generation (n = 20), or protein translation (n = 20). Using motif discovery tools revealed that genes in the contractile module were found to contain a TATA-box followed by a CACC-box, and are depleted in other GC-rich motifs; whereas genes in the translation module contain a pyrimidine-rich initiator, Elk-1, SP-1, and a novel motif with a GCGC core. Using a naïve Bayes classifier revealed that patterns of motifs are statistically predictive of expression patterns, with odds ratios of 2.7 (contractile), 1.9 (energy generation), and 5.5 (protein translation). Conclusion We identified patterns comprised of putative <it>cis</it>-regulatory motifs enriched in the upstream promoter sequence of genes that undergo similar changes in expression secondary to cardiomyopathies of various etiologies. Our analysis is a first step towards understanding transcription factor networks that are active in regulating gene expression during degenerative heart disease.</p

Comparative analysis of cis-regulation following stroke and seizures in subspaces of conserved eigensystems

Public Library of Science (PLOS)

Probabilistic Inference of Transcription Factor Binding from Multiple Data Sources

Author: A Ambesi-Impiombato
A Bernard
A Beyer
A Sandelin
A Sandelin
A Siepel
AFA Smit
Alistair G. Rust
B Ren
CE Lawrence
CL Warren
CP Robert
CT Harbison
D GuhaThakurta
D Husmeier
D Husmeier
David Jones
DB Gordon
DJ Reiss
DJ Wilkinson
DT Holloway
DT Holloway
E Blanco
E Segal
E Segal
E Wingender
EH Davidson
G Chen
G Thijs
G Thijs
GD Stormo
GE Crawford
H Huang
H Lähdesmäki
H Steck
Harri Lähdesmäki
Ilya Shmulevich
IV Bajić
J Taylor
JD Hughes
JM Claverie
K Quandt
K Thomas
KD MacIsaac
KP Murphy
L Hertzberg
L Narlikar
L Narlikar
L Narlikar
L Zhang
M Eisenstein
M Kellis
M Levine
M Tompa
MA Beer
MC Frith
MF Berger
MJL de Hoon
ML Bulyk
N Friedman
N Rajewsky
ND Heintzman
O Hallikas
OV Kel-Margoulis
Q Zhou
R Siddharthan
R Staden
S Cawley
S Mukherjee
S Sinha
S Sinha
SB Montgomery
SJ Maerkl
SP Brooks
ST Jensen
T Chen
T Fawcett
T Reguly
TD Wu
TI Lee
TL Bailey
TL Bailey
VD Marinescu
W Pan
WJ Kent
WP Lehrach
WW Wasserman
X Liu
X Xie
XS Liu
Y Barash
Y Barash
Y Qi
Y Tamada
Publication venue: Public Library of Science
Publication date: 01/03/2008
Field of study

An important problem in molecular biology is to build a complete understanding of transcriptional regulatory processes in the cell. We have developed a flexible, probabilistic framework to predict TF binding from multiple data sources that differs from the standard hypothesis testing (scanning) methods in several ways. Our probabilistic modeling framework estimates the probability of binding and, thus, naturally reflects our degree of belief in binding. Probabilistic modeling also allows for easy and systematic integration of our binding predictions into other probabilistic modeling methods, such as expression-based gene network inference. The method answers the question of whether the whole analyzed promoter has a binding site, but can also be extended to estimate the binding probability at each nucleotide position. Further, we introduce an extension to model combinatorial regulation by several TFs. Most importantly, the proposed methods can make principled probabilistic inference from multiple evidence sources, such as, multiple statistical models (motifs) of the TFs, evolutionary conservation, regulatory potential, CpG islands, nucleosome positioning, DNase hypersensitive sites, ChIP-chip binding segments and other (prior) sequence-based biological knowledge. We developed both a likelihood and a Bayesian method, where the latter is implemented with a Markov chain Monte Carlo algorithm. Results on a carefully constructed test set from the mouse genome demonstrate that principled data fusion can significantly improve the performance of TF binding prediction methods. We also applied the probabilistic modeling framework to all promoters in the mouse genome and the results indicate a sparse connectivity between transcriptional regulators and their target promoters. To facilitate analysis of other sequences and additional data, we have developed an on-line web tool, ProbTF, which implements our probabilistic TF binding prediction method using multiple data sources. Test data set, a web tool, source codes and supplementary data are available at: http://www.probtf.org

Assessing Computational Methods of Cis-Regulatory Module Prediction

Author: A Bruhat
A Siepel
A Sosinsky
A Visel
AB Rose
AG Clark
AL Halpern
AM Moses
B Prud'homme
B Shi
BK Peterson
BP Berman
BY Chan
Christina Leslie
CM Bergman
CM Bergman
D Kolbe
D Papatsenko
DA Kleinjan
DC King
DC King
DE Schones
DM Jeziorska
DS Johnson
E Birney
E Davidson
E Emberly
E Segal
E Wingender
G Bejerano
GM Euskirchen
H Wang
H Weintraub
JB Warner
Jing Su
JL Kabat
JR Stone
JS Jakobsen
KH Surinya
KJ Won
L Li
LP Lim
M Bieda
M Blanchette
M Brudno
M Hasegawa
MC Frith
MD Schroeder
MD Wilson
MS Halfon
MS Halfon
MZ Ludwig
N Bray
N Ghanem
N Gompel
N Pierstorff
ND Heintzman
ND Heintzman
O Hallikas
O Johansson
OV Kel-Margoulis
P Van Loo
PC FitzGerald
PJ Sabo
Q Zhou
Q Zhou
R Godbout
RP Zinzen
S Aerts
S Aerts
S Batzoglou
S Karlin
S MacArthur
S Richards
S Sinha
S Sinha
S Sinha
Sarah A. Teichmann
SC Parker
SE Celniker
T Sandmann
T Strachan
T Waleev
Thomas A. Down
TL Bailey
TM Williams
V Ferretti
V Gotea
W Krivan
WW Wasserman
X He
X He
XY Li
Publication venue: Public Library of Science
Publication date: 01/01/2010
Field of study

Computational methods attempting to identify instances of cis-regulatory modules (CRMs) in the genome face a challenging problem of searching for potentially interacting transcription factor binding sites while knowledge of the specific interactions involved remains limited. Without a comprehensive comparison of their performance, the reliability and accuracy of these tools remains unclear. Faced with a large number of different tools that address this problem, we summarized and categorized them based on search strategy and input data requirements. Twelve representative methods were chosen and applied to predict CRMs from the Drosophila CRM database REDfly, and across the human ENCODE regions. Our results show that the optimal choice of method varies depending on species and composition of the sequences in question. When discriminating CRMs from non-coding regions, those methods considering evolutionary conservation have a stronger predictive power than methods designed to be run on a single genome. Different CRM representations and search strategies rely on different CRM properties, and different methods can complement one another. For example, some favour homotypical clusters of binding sites, while others perform best on short CRMs. Furthermore, most methods appear to be sensitive to the composition and structure of the genome to which they are applied. We analyze the principal features that distinguish the methods that performed well, identify weaknesses leading to poor performance, and provide a guide for users. We also propose key considerations for the development and evaluation of future CRM-prediction methods

CiteSeerX

Helmholtz Zentrum für Infektionsforschung Repository

MatrixCatch - a novel tool for the recognition of composite regulatory elements in promoters

Author: A Kel
Alexander E Kel
DK Biswas
E Jacox
E Shelest
E Sterneck
Edgar Wingender
Elena V Deineko
Igor V Deyneko
K Klepper
M Xu
M Ye
MH Kim
MH Kim
MI Diamond
Olga V Kel-Margoulis
OV Kel-Margoulis
P Van Loo
Siegfried Weiss
T Heinemeyer
T Waleev
V Matys
W Krivan
WG Butscher
WW Wasserman
Publication venue: 'Springer Science and Business Media LLC'
Publication date: 01/01/2013
Field of study

Accurate recognition of regulatory elements in promoters is an essential prerequisite for understanding the mechanisms of gene regulation at the level of transcription. Composite regulatory elements represent a particular type of such transcriptional regulatory elements consisting of pairs of individual DNA motifs. In contrast to the present approach, most available recognition techniques are based purely on statistical evaluation of the occurrence of single motifs. Such methods are limited in application, since the accuracy of recognition is greatly dependent on the size and quality of the sequence dataset. Methods that exploit available knowledge and have broad applicability are evidently needed.We developed a novel method to identify composite regulatory elements in promoters using a library of known examples. In depth investigation of regularities encoded in known composite elements allowed us to introduce a new characteristic measure and to improve the specificity compared with other methods. Tests on an established benchmark and real genomic data show that our method outperforms other available methods based either on known examples or statistical evaluations. In addition to better recognition, a practical advantage of this method is first the ability to detect a high number of different types of composite elements, and second direct biological interpretation of the identified results. The program is available at http://gnaweb.helmholtz-hzi.de/cgi-bin/MCatch/MatrixCatch.pl and includes an option to extend the provided library by user supplied data.The novel algorithm for the identification of composite regulatory elements presented in this paper was proved to be superior to existing methods. Its application to tissue specific promoters identified several highly specific composite elements with relevance to their biological function. This approach together with other methods will further advance the understanding of transcriptional regulation of genes.peerReviewe